Journal of Vision
● Association for Research in Vision and Ophthalmology (ARVO)
Preprints posted in the last 30 days, ranked by how well they match Journal of Vision's content profile, based on 110 papers previously published here. The average preprint has a 0.06% match score for this journal, so anything above that is already an above-average fit.
de Jong, J.; Sergent, C.; Wexler, M.
Show abstract
The temporal resolution of vision is seriously limited. However, the response to a very brief flash, called the impulse response, is already quite sluggish at the earliest stages of vision, potentially obscuring the true temporal resolution of the rest of the visual system. Faster monitors that produce briefer flashes are subject to diminishing returns because, by definition, they cannot elicit responses that are any briefer than the impulse response. Here, taking inspiration from previous attempts, we develop a novel technique for presenting flashes that elicit 'briefer-than-brief' visual responses. Using a simple deconvolution technique, we reverse-engineer the visual response and estimate the form that the stimulus should take to elicit the response that a faster visual system would produce to a normal flash. Using psychophysics on human observers, we demonstrate that these 'briefer-than-brief' (BTB) flashes partially bypass the temporal limits presumably imposed by the early visual system using two paradigms: one that requires temporal segregation and one that requires temporal integration of sequential flashes. With BTB flashes, human observers successfully isolated two successive flashes at shorter intervals than with conventional flashes, improving temporal resolution by around 16%. We found that BTB stimuli not only improved temporal resolution, but also induced poorer performance on tasks requiring temporal integration, suggesting that the visual responses elicited by BTB flashes overlap less in time due to their briefer duration. In sum, our findings suggest that, using reverse-engineered stimuli, we can alleviate a temporal bottleneck that probably originates from the earliest stages of vision. In doing so, we allow higher visual areas to operate at a higher temporal resolution than previously thought possible.
Mendez, A. H.; Otero-Millan, J.; de la Malla, C.; Lopez-Moliner, J.
Show abstract
Rigorously tracking eye and head behavior in space is key to building realistic models of the stimulus that reaches our retina. The motion structure of this stimulus or retinal flow - the substrate for self and object motion processing - is created by the relative movement of the eyes with respect to the world. Characterizing this stimulus requires tracking the eyes three degrees of freedom in the head and the heads six degrees of freedom in the world. While vertical and horizontal eye rotations have been described during locomotion in the context of gaze stabilization (Moore et al, 2001), the component around the line of sight - torsion - has remained difficult to quantify, and how all three rotational components jointly contribute to retinal flow during self-motion remains largely unexplored. Here, we leveraged head-mounted technology to estimate eye torsion in ten subjects as they walked towards a distant target in a fast and slow condition (from 14 to 4 meters away from the target, see Fig. 1A). More specifically, we combined automatic feature tracking with gaze-constrained simulations of eye rotations and camera projection to recover torsion from image data. We then estimated flow curl in head and retina centered frames in two scenarios: torsion as estimated from our data and with no torsion. We show that the eyes torsional component compensates for the roll component of heads angular displacement, altering the incoming visual flow in ways that are relevant for the extraction of self-motion parameters from retinal flow. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=141 SRC="FIGDIR/small/743586v1_fig1.gif" ALT="Figure 1"> View larger version (38K): org.highwire.dtl.DTLVardef@391b9eorg.highwire.dtl.DTLVardef@1444510org.highwire.dtl.DTLVardef@1121e16org.highwire.dtl.DTLVardef@754b6a_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFig 1.C_FLOATNO A. Top. Custom-made head-mounted device combining the Neon eye tracker (Pupil Labs), an RGB camera and a dimmable light. Bottom. Four 3D frames of reference (FoR) are relevant for this study, two static (world and locomotion) and two subject centered (head and eye). The Z axis of the locomotion, head and eye FoRs are approximately aligned throughout the trial. For the locomotion FoR the Z axis is fixed in the world and points forward (towards the target). The heads Z axis moves with the head but - as subjects are fixating a target along their path -, it also points approximately forward. The eyes Z axis also moves with the head and its exact forward orientation will depend on compensatory eye movements. B. Left. Blue dots represent the Z component of the heads orientation vector (on the locomotion frame) on the X axis, and the sum of all three components on the Y axis; for each frame for all corpus data. Blue contour is the 75th percentile 2D density distribution of the blue dots. Red and violet contours represent the 75th percentile for the X and Y components of head orientation, respectively. Right. Same logic but applied to the heads velocity vector. C. Left. Two examples showing the mean rotation of iris features over the course of a slow (top) and fast (bottom) trial. Colored lines show each of the 561 simulated cameras for a given scenario (one color per scenario); the black line shows the camera from the empirical data. Right. Trial-level mean fit score of each scenario with the empirical data is represented as a function of each subjects fitted gain. 20 dots represent 10 subjects x 2 trials. On the rightmost column, all values are aligned vertically to show the mean fit score across trials for the three scenarios. Size indicates the 75th percentile of head z component for each trial. C_FIG
Spitschan, M.
Show abstract
PurposePupil diameter in daily life depends on both the light reaching the eye and the observers age, but established prediction formulas require laboratory quantities that are rarely measured in natural environments. We developed a compact age-corrected model that predicts pupil diameter from melanopic equivalent daylight illuminance (mEDI). MethodsWe used an existing field dataset in which binocular pupil diameter and near-corneal spectral irradiance were recorded while 83 adults aged 18-87 years moved through indoor and outdoor environments. The analysis included 10,082 valid paired observations. We fitted a bounded sigmoid relating pupil diameter to mEDI and age, with each participant given equal influence, and assessed prediction in participants excluded from model fitting. Performance was compared with simpler models, a flexible generalised additive model (GAM), and Watson-Yellott predictions based on assumed field geometry. ResultsPupil diameter decreased smoothly as mEDI increased. Age primarily reduced the difference between pupils in dim and bright conditions, by 0.768 mm per decade, while the predicted bright-light diameter changed little with age. In held-out participants, the bounded model had a participant-balanced root mean squared error (RMSE) of 0.630 mm and mean absolute error of 0.537 mm. The GAM had a slightly lower point-estimate RMSE of 0.610 mm, but the difference was small and uncertain. The bounded model outperformed the tested log-linear, reduced, age-only, and Watson-Yellott alternatives. ConclusionAge and mEDI are sufficient to provide useful population-average pupil predictions across the observed adult age and real-world light range. The model is transparent, physiologically bounded, and nearly as accurate as a flexible GAM, but predictions approaching darkness remain uncertain because valid mEDI measurements were not available in that range. Key pointsO_LIA compact equation predicts population-average pupil diameter from age and mEDI alone. C_LIO_LIAge mainly compresses the pupils response range by reducing pupil diameter under dimmer conditions. C_LIO_LIPrediction error in unseen participants was close to that of a flexible GAM, without requiring a fitted smooth object. C_LIO_LIThe model is intended for the observed adult age and field-light range, not for extrapolation into darkness. C_LI
Pandey, P.; Pethe, S. R.; Indrajeet, I.; Ray, S.
Show abstract
Introduction: Decision making for selecting an object or a course of action from possible alternatives largely depends on our perceptual ability modulated by attention. When multiple stimuli appear close together in time, processing one stimulus can temporarily impair the processing of another due to temporal limitations of attention. Observers frequently fail to detect the second target (T2) presented within a few hundred milliseconds after the first target (T1) in a stream of stimuli, which is commonly known as attentional blink (AB). Existing theories attribute this perceptual lapse to T1 processing, distractor interference, or transient attentional gating; however, the computations underlying suppressive mechanism remains unresolved. We investigated whether pupil-size could reveal the underlying mechanisms of AB and predict conscious perception on a trial-by-trial basis. Methods: Pupil diameter and gaze locations were recorded using an infrared eye tracker. Machine learning techniques were used to classify trials when T2 was detected versus when it was not, after correct identification of T1, during an AB task from the pupil dynamics, which also yielded attentional episode (AE) associated with each element in the stream of visual stimuli when deconvolved. Results: Cross-validating classifiers achieved near-perfect accuracy not only in distinguishing but also predicting perceptual outcomes on a single-trial basis. AEs exhibited greater power when T2 was detected than when it was missed; the differential power in AEs on a logarithmic scale was highly synced with the differential pupil size. Conclusions: Collectively, these findings establish a framework for predicting attention-driven perceptual outcomes from pupil-dynamics at finer time-scale.
Yeh, L.-C.; Kaiser, D.
Show abstract
The attentional blink is a well-known phenomenon illustrating the limitations of human attention: When two visual targets are presented in rapid succession, identification of the second target is often impaired. While the attentional blink is known to attenuate when targets share perceptual features or category membership, real-world objects are also linked through contextual associations, shaped by objects typically occurring within the same environments. Here, we devised an attentional blink experiment in which we orthogonally manipulated contextual and categorical relationships between the two targets while controlling for their perceptual similarity. As the key result, contextual associations facilitated identification of the second target but impaired identification of the first target. These findings suggest that contextual associations yield distinct benefits and costs for visual cognition, where enhanced attentional access to subsequent targets is traded off against increased interference between targets.
Vlachou, M. E.; Thomas, E.; Blouin, J.
Show abstract
In this paper, we address the problem of quantifying similarity between planar 2D shapes, which is relevant to studies of internal representations in cognitive, developmental, and neurological research. We designed a set of test shapes arranged along a visually defined perceptual similarity gradient and used them to evaluate classical geometric methods for shape comparison, including Procrustes and Chamfer distance, as well as a convolutional neural network (CNN)-inspired feature-based method. Based on the limitations identified for these individual methods, we developed a hybrid Geometric-Feature Similarity (GFS) algorithm that combines geometric alignment, global contour properties, and convolutional feature-based descriptors into a unified weighted similarity score. By combining global geometric information with local structural features, the GFS algorithm more accurately reproduces human perceptual judgments of shape similarity than either geometric or feature-based methods alone. Requiring neither network training nor large labelled datasets, the proposed algorithm provides an efficient and interpretable tool for a broad range of studies involving quantitative shape comparison.
Menetrey, M. Q.; Pascucci, D.
Show abstract
Several theories propose that perception and attention are governed by rhythmic processes that give rise to periodic fluctuations in behavior. However, empirical support for behavioral rhythms has been derived largely from paradigms involving brief, static stimuli. Here, we introduce a temporal averaging task requiring integration of rapidly unfolding visual features. Across three experiments, we tested averaging of orientation, size, and color under different eccentricity conditions. We used a temporally weighted averaging model to assess whether the influence of individual stimulus samples on perceptual estimates exhibits periodic modulation over time. We found no common rhythmic signature across tasks. Instead, orientation and size judgments showed reliable low-frequency modulations (<2.5 Hz), whereas color judgments showed only weak trends. Higher-frequency components (~3.5-8 Hz), often linked to theta and alpha rhythms, were observed only in a subset of participants and were limited to parafoveal orientation processing. These findings challenge the notion of universal behavioral rhythms and instead suggest that temporal dynamics are task-dependent, with slow oscillatory processes emerging as the most consistent feature.
Dahech, H.; Minami, T.; Nakauchi, S.; Tamura, H.
Show abstract
Why does an angry face feel uncomfortable? The answer is that it signals a threat. However, a face is only part of an encounter, and distance, facial stimulus type, and gaze may shape discomfort regardless of perceived anger. To separate these cues, we conducted three within-subjects virtual reality experiments. In each experiment, 24 adults viewed avatars at intimate, personal, and social distances (30, 100, and 300 cm, respectively) and rated the faces perceived anger and their own discomfort; head movement was recorded in Experiments 2 and 3. In Experiment 1, the expression (angry, neutral) and facial color (natural, red) were crossed with distance; in Experiment 2, a featureless mannequin served as a nonface comparison; and in Experiment 3, the gaze direction (direct, averted) was manipulated. Expression primarily determined perceived anger, whereas distance predominantly determined discomfort: A nearby neutral face was uncomfortable despite low perceived anger (Experiment 1). A neutral human face was more uncomfortable than a mannequin, although both received similarly low perceived-anger ratings (Experiment 2). Direct gaze increased the discomfort without changing perceived anger (Experiment 3). Backward head movement exhibited a similar pattern, with participants leaning back more from human faces than from the mannequin. These results indicate that the discomfort associated with an angry face is not merely explained by perceived anger. Instead, social discomfort was differentially associated with interpersonal distance and gaze direction and differed between the human-face and mannequin conditions.
Yildiran, O. F.; Ni, L.; Landy, M. S.
Show abstract
Previous work showed that observers integrate audiovisual duration cues optimally when cue-conflict is small. Does causal inference lead to a breakdown of audiovisual integration when duration conflicts are large? We addressed this by testing a wide range of duration cue-conflicts. Participants compared the auditory durations of a test and a standard stimulus. Audiovisual durations were consistent in the test stimulus, but differed by seven conflict durations (up to 250 ms) in the standard. Two levels of auditory noise were tested. Auditory duration percepts shifted systematically toward the visual duration, especially with high auditory noise. The shift was proportional to cue-conflict magnitude, inconsistent with causal inference. We compared several models. A heuristic model in which the observer probabilistically switches between the visual and auditory cues was preferred for most participants, although performance differences across models were small. Within the tested conflict range, the forced fusion, causal inference, and probabilistic cue switching models produced overlapping, near-linear shifts as a function of cue-conflict. Model simulations further revealed that given the measured sensory noise, forced fusion and causal inference can be discriminated only with unreasonably large conflicts. Together, while our results suggest that observers do not rely on causal inference when judging auditory durations under our conditions, high sensory encoding noise in auditory duration limits the discriminability of competing computational models.
Darjani, N.; Bakhtiari, S.; Vaziri-Pashkam, M.; Robert, S.
Show abstract
The human visual system integrates both static and dynamic information to support form and shape perception, yet the computational principles underlying the integration of motion for object recognition remain unclear. Artificial neural networks (ANNs) offer a computational framework for developing and testing hypotheses about these principles: if ANNs trained on motion-related tasks develop representations that align with brain activity and support object categorization, this would suggest that the training objectives and architectural constraints of these networks may capture key aspects of motion processing in biological visual systems in general, and motion processing for object recognition, in particular. Here, we investigated this question using "object kinematograms", stimuli in which object form is conveyed solely through motion cues. We measured neural responses of two higher regions of the lateral and the dorsal visual pathways, respectively, with strong sensitivity to dynamic cues from objects: lateral occipitotemporal cortex (LOTbio), and left supramarginal gyrus (SMGlh), as well as primary visual cortex (V1). We compared brain responses to representations extracted from two neural networks: SlowFast, a dual-pathway architecture trained on action recognition that processes slow- and fast-varying visual information with cross-pathway integration, and DorsalNet, a model of the primate dorsal visual pathway trained on embodied self-motion estimation. Representational similarity analysis revealed distinct representational profiles across brain areas, demonstrating functional specialization in motion-based form processing. LOTbio was best characterized by the slow pathway of the SlowFast model, whereas SMGlh showed strong similarity to both models. Critically, we found that representations aligned with brain activity also better supported behavioral function: the full SlowFast model, incorporating both slow and fast pathways, outperformed other models in few-shot categorization of object kinematograms and showed the highest similarity to human perceptual judgments. These findings demonstrate that with appropriate inductive biases, specifically, dual-pathway architectures for multi-scale motion processing and training objectives focused on dynamic visual tasks, ANNs can develop functionally useful representations of motion-defined forms that exhibit better alignment with the visual regions involved in processing dynamic visual signals.
Chan, A. Y. C.; Shimojo, S.
Show abstract
This study characterizes how people combine visual and tactile directional cues while acting in a fully immersive 360{degrees} virtual environment. Participants used a vibrotactile belt and VR headset to localize targets while we manipulated visual reliability and the spatial discrepancy between visual and tactile signals. Behaviorally, degraded visual input made visual responses slower, less precise, and more susceptible to tactile pull, whereas tactile-guided responses remained comparatively stable. We then asked whether these behavioral changes reflected a change in multisensory binding or a change in sensory uncertainty. A Bayesian Causal Inference (BCI) framework captured the structure of behavior under high visual reliability and continued to track individual differences under low visual reliability, even though its absolute goodness-of-fit decreased. Under extreme visual noise, Bayesian Information Criterion sometimes favored a simpler Maximum Likelihood Estimation (MLE) model, but MLE showed poor absolute fit and did not capture meaningful behavioral variability. This dissociation shows that statistical parsimony and explanatory validity can diverge when behavior becomes highly variable. BCI-derived parameters further indicated that degraded vision increased visual uncertainty, while the prior tendency to bind visual and tactile cues remained stable. Kinematic analyses added a complementary insight: early movement trajectories were strongly shaped by tactile signals, even when final localization was visually guided. Together, these findings suggest that visual-tactile integration in 360{degrees} environments depends on sensory reliability and task demands, with tactile cues providing fast body-centered guidance when visual information is limited.
Sklyar, Y.; Hendler, S.; Schonberg, T.
Show abstract
Museum visits typically follow curator-defined routes that constrain how visitors shape their own experience, yet choice is widely held to heighten engagement, autonomy, and enjoyment. Virtual reality (VR) offers a setting in which to study these processes because it combines ecological immersion with precise, continuous behavioral measurement. We investigated (i) whether VR- derived behavioral signals are associated with self-reported enjoyment during a virtual museum tour, and (ii) whether the level of agency afforded to visitors influences enjoyment. Forty-eight adults completed a room-scale, life-size VR tour (8 * 4 m) of seven paintings from the Tel Aviv Museum of Art, each accompanied by a synchronized audio guide. Synchronized gaze and head- position streams were logged continuously (50 Hz) and segmented into painting-level viewing episodes using a trial-and-tile pipeline that intersects each painting's trial interval with an empirically defined spatial window in front of the canvas. Participants were randomly assigned to one of three agency conditions, Active (choice before every artwork), Semi-Active (choice for the first three), or Passive (fixed route),while the artwork sequence was held identical. Self- reported enjoyment at the tour and painting levels did not differ reliably across agency conditions. Among VR-derived measures, gaze engagement during the audio guide showed the clearest (though modest) association with painting-level liking, whereas locomotion and pacing measures were weak and inconsistent predictors. Agency nonetheless reliably modulated several gaze- and time-based viewing measures. The findings reveal a dissociation between subjective enjoyment and the micro-structure of viewing, and establish a reusable framework for full-tour, painting-level behavioral analysis in immersive settings.
Lin, C.-H. S.; Terence, N.; Garrido, M.
Show abstract
Bayesian decision theory proposes that people make statistically rational decisions by combining prior knowledge with sensory information (likelihoods). This framework successfully explains many aspects of human behaviour. However, debate persists over whether people perform precise Bayesian computations (i.e., explicit Bayesian strategy) or rely on less demanding strategies - such as approximations or heuristics - that produce Bayesian-like behaviour (i.e., implicit Bayesian strategy). To address this, we examined people's sensitivity to metamers: different prior-likelihood combinations yielding identical optimal policies. An explicit Bayesian observer would show a temporary performance drop immediately after a switch of prior-likelihood combination, followed by recovery, reflecting prior updating. In two studies, we trained participants to estimate hidden target locations drawn from a Gaussian prior. On each trial, scattered dots provided likelihood information. Over time, participants learned the prior and combined it with likelihood information to infer target locations. We then covertly introduced an untrained prior-likelihood metamer. Unlike explicit Bayesian observers, participants' performance declined after the switch and persisted throughout the untrained pair presentation. This finding challenges strict Bayesian interpretations of task performance and suggests that participants rely instead on likelihood-sensitive strategy that is neither explicit Bayesian nor does it not fully integrate prior information. Our study demonstrates how metamer manipulations can distinguish behaviour that merely appears Bayesian, from behaviour genuinely produced by Bayesian computations, and calls for the use of metamers for ruling out alternative explanations of Bayesian-like behaviours.
Ha, L.; Sun, C.; Tang, R.
Show abstract
Analysis does not always enhance aesthetic experience. Philosophical accounts have long suggested that decomposing an aesthetic experience into determinate components may weaken it, yet this possibility has rarely been tested experimentally. To examine whether, when, and how analysis produces divergent effects on aesthetic experience, we conducted two experiments manipulating analysis depth. Experiment 1 showed that, during affective analysis of visual art, deep analysis produced a significantly weaker increase in aesthetic ratings than shallow analysis. In Experiment 2, we selected this condition to investigate the underlying mechanism. The behavioral effect was replicated: deep analysis removed the increase produced by shallow analysis without reducing ratings below the image baseline. Frequency-resolved brain network analysis further revealed a stronger task-related component and higher spatial entropy within the default mode network under deep analysis. Network-behavior correlations observed under shallow analysis were absent under deep analysis, suggesting reduced correspondence between the default-mode network (DMN) organization and aesthetic experience. Exploratory analyses further showed that spatial weights in the lateral temporal cortex and inferior parietal lobule were associated with smaller increases in aesthetic ratings. Together, these findings indicate that deeper analysis can selectively weaken improvements in aesthetic experience by altering how affective information is organized within the DMN.
Flieger, P.; Stecher, R.; Kaiser, D.
Show abstract
Humans rapidly assess the beauty of natural scene images. Previous EEG work suggests neural representations of beauty emerge early and are temporally sustained. Complementary fMRI work pinpoints the neural correlates of beauty to visual, frontal, and default-mode network areas. An integrated view of the spatiotemporal dynamics that give rise to the perception of beauty, however, is lacking. Beyond the beauty of the depicted scene, the quality of the image itself influences its perceived beauty, and it is unknown how the brain separates these two factors. To address these questions, we recorded EEG (N = 52) and fMRI (N = 29) data while participants rated the beauty of 100 natural scene photographs. Another group of participants (N = 46) rated the image quality of the same photographs. Separate representational similarity analyses on the EEG and fMRI data revealed early and sustained beauty-related representations across widespread cortical areas. In contrast, representations of image quality emerged earlier, had markedly different representational dynamics, and were predominantly localized to visual cortex. In a model-based EEG-fMRI fusion analysis, we investigated how the correspondence between temporally resolved EEG signals and spatially resolved fMRI signals is explained by beauty ratings. Our results suggest that beauty-related representations emerge early (from around 165ms and peaking at 275ms post-onset), are long-lasting, and primarily originate from high-level visual cortex. This spatiotemporal signature persisted when controlling for image-quality ratings. Our findings emphasize the importance of perceptual processing for perceived beauty and suggest that the brain represents aesthetic appeal independently of image quality.
Arora, K.; Gayet, S.; Kenemans, L.; Naber, M.; Chota, S.; Van der Stigchel, S.
Show abstract
Attention forms a key, yet elusive component of visual processing. We shift attention constantly across the visual field to enhance processing of relevant locations or stimuli in service of goal-directed behavior. Here, we investigate a fundamental property of attentional shifts: when covertly shifting from one location to another, does visual attention "travel" (enhance processing at intermediate locations) or "teleport" (not interact with intermediate locations)? While most shifts of information or movement "travel" (e.g., eye and body movements, neuronal signalling), for attentional shifts such intermediate processing enhancement might not be necessary, nor functional. To answer this question while tackling the difficulty of tracking covert attentional dynamics, we conducted an EEG-eyetracking experiment (n=24) paired with Rapid Invisible Frequency Tagging (RIFT). This let us track attentional enhancement during covert attentional shifts with high temporal and spatial precision. We successfully registered covert shifts, but found no evidence for any attentional modulation in-between the start and end-point of an attentional shift. Additional insilico modelling of attentional shifts confirmed that our method was sensitive enough to pick up on these modulations had they been present. Our results support a "teleporting" model of attention, suggesting that attention is implemented in a fundamentally different manner compared to overt visual behaviours such as eye movements.
Heirani Moghaddam, S.; Decarie, A.; Chua, R.; Cressman, E. K.
Show abstract
In the current experiment, we compared reported perceptual awareness of the visuomotor rotation to motor awareness of changes in reaches established using the process dissociation procedure and drawing task following visuomotor adaptation to a large (50 degrees; R50 group) or a small (30 degrees; R30 group) cursor rotation. Results revealed that perceptual and motor awareness did not differ in magnitude for the R50 group and were significantly correlated. In contrast, while the R30 group perceptually reported being aware of the visuomotor rotation, motor awareness was significantly less and responses were not significantly correlated across tasks. Overall, results suggest that perceptual and motor tasks assess different processes underlying visuomotor adaptation to a small cursor rotation, such that perceptual awareness of the visuomotor rotation is not reflected in reaching performance on tasks assessing motor awareness.
Le Moël, F.; Webb, B.
Show abstract
Insects solve complex behavioural tasks with remarkable efficiency, using minimal neural hardware tuned to the specific requirements of their ecological niches. To truly understand or replicate these behaviours, it is insufficient to model the brain in isolation: one must account for the dynamic, closed-loop interactions between the environment, the physical organisation of the sensory periphery, and internal biophysical dynamics. To address these issues for visually controlled behaviours, we present RhabdoForge, a modular, hardware-agnostic and high-performance rendering framework specifically designed for insect neuroethology and neuromorphic research. Designed for seamless integration into Python-based workflows, RhabdoForge implements both real-time ray-tracing and stochastic path-tracing using hardware-agnostic GPU pipelines. Crucially, the engine moves beyond the static "ommatidium-as-a-pixel" paradigm by introducing a fully parametrisable model where every layer of the compound eye (from the geometric shape and the topological lattice to the internal rhabdomere blueprint) is a discrete, swappable component. The engine is capable of simulating the high-frequency, sub-ommatidial rhabdomere photomechanical actuation, allowing for the investigation of a variety of active sensing phenomena within a real-time closed-loop environment. The framework also includes an automated morphological pipeline that allows transforming 2D anatomical data into faithful 3D sensory models. We validate the engine through two case studies: a closed-loop optic-flow centring response in a virtual tunnel, and the recovery of spatial hyperacuity via rhabdomere microsaccades. By providing a bridge between high-fidelity visual ecology and neuromorphic modelling, RhabdoForge enables researchers to explore how the interplay of sensory optics and neural processing can generate complex behaviour in both biological and artificial agents.
Sharifi Nowghabi, A.; Sharghilavan, S.; Bagheri, A.; Izadifar, M.
Show abstract
Wayfinding in hospitals is often hindered by ineffective signage; however, the cognitive mechanisms of healthcare wayfinding symbols comprehension remain under-researched. This study utilized eye-tracking and spatial gaze mapping to examine how visual complexity, abstraction, and human figuration modulate perception in 40 healthy adults viewing 24 hospital-related healthcare wayfinding symbols. Results indicate that pupil size is a sensitive physiological marker of cognitive load, significantly influenced by visual complexity ({chi}2 = 11.32, p = .022) and abstraction ({chi}2 = 7.49, p = .027). Human figuration reduced fixation duration and increased saccade amplitude, facilitating efficient semantic integration. Furthermore, human-centric healthcare wayfinding symbols elicited streamlined gaze trajectories, whereas abstract/complex designs induced chaotic scanpaths. These findings suggest that human figuration acts as a cognitive scaffold, reducing mental effort. We provide evidence-based guidelines for optimizing healthcare wayfinding symbols by prioritizing human body representations and balancing abstraction levels. HighlightO_LIPupil size indexes cognitive load during symbol comprehension. C_LIO_LIHuman figuration cuts fixation duration, boosting wayfinding efficiency. C_LIO_LIAbstract symbols increase pupil dilation, raising cognitive load. C_LI
Prince, J. S.; Wang, B.; Fel, T.; Jagadeesh, A. V.; Vaziri, P. A.; Alvarez, G. A.; Livingstone, M. S.; Konkle, T.
Show abstract
Leading deep neural network encoding models predict visual cortical responses with nearly indistinguishable accuracy, raising the strong inference that these models have converged on the same underlying brain-aligned parameterization of natural image space. Here we demonstrate that this is not the case. We introduce axis-aligned feature accentuation, which converts each model's fitted encoding axis into graded stimulus perturbations that are predicted to parametrically control neural firing within and beyond the natural-image range. We generated over 27,500 controller stimuli from ten leading vision models and presented them to five macaques in closed-loop experiments targeting early, mid-, and high-level visual areas. Despite matched natural image predictivity, models diverged strongly in their ability to control neural firing using accentuated stimuli, revealing that most model encoding axes failed to capture the precise tuning of their corresponding neurons. The two adversarially trained models showed a consistent advantage, though adversarial robustness was only weakly predictive of neural control across other models. Instead, control was better predicted by the spatial frequency structure of the input gradient: the distribution of pixels influencing each encoding axis. Overall, these results establish neural control via axis-aligned feature accentuation as a causal method to assess the alignment between how neurons and models parameterize the visual world.